Papers with computational methods
Copied to clipboard
| Challenge: | a tutorial focuses on computational models for conversational structure, summarization and sentiment detection, and group dynamics. |
| Approach: | a tutorial will provide examples of specific NLP tasks for conversational structure, summarization and sentiment detection, and group dynamics. |
| Outcome: | The tutorial focuses on the three areas of conversational structure, summarization and sentiment detection, and group dynamics. |
Copied to clipboard
| Challenge: | Argumentation is a rhetorical device that asserts propositions implicitly, but few studies have examined the issue. |
| Approach: | They propose a computational method for extracting propositions that are implicitly asserted in questions, reported speech, and imperatives in argumentation. |
| Outcome: | The proposed models are based on a corpus of 2016 debates and online commentary. |
Copied to clipboard
| Challenge: | Political scientists have developed and adopted natural language processing (NLP) methods to exploit text as an additional source of data in their analyses. |
| Approach: | This tutorial aims to provide a gentle introduction to methods and tasks related to computational analysis of political texts from both communities. |
| Outcome: | The main goal of this tutorial is to bring the two research communities closer to each other and contribute to faster and more significant developments in this interdisciplinary area. |
Copied to clipboard
| Challenge: | PD is the second most common neurodegenerative disorder after Alzheimers disease . speech impairments are one of the earliest manifestations in PD patients . |
| Approach: | They propose to analyze the speech signals of PD patients and healthy control subjects in three different languages: German, Spanish, and Czech. |
| Outcome: | The proposed model can discriminate between PD patients and HC subjects even when the language used for train and test is different. |
Copied to clipboard
| Challenge: | 'agency' is the freedom and capacity of an entity to act, and the corresponding Natural Language Processing (NLP) task involves automatically detecting attributions of agency to entities in text. |
| Approach: | They propose a schema to annotate a dataset for agency attribution and formulate additional research questions by applying NLP models. |
| Outcome: | The proposed framework draws on semantic frame analysis, role labelling and related techniques. |
Copied to clipboard
| Challenge: | Logical metonymies are type clashes between an event-selecting verb and an entity-denoting noun . they are typically interpreted by inferring a hidden event on the basis of contextual cues . |
| Approach: | They propose to use probabilistic and distributional models to model logical metonymy interpretation . they compare models with the best Transformer-based models and some traditional distributional ones . |
| Outcome: | The proposed models perform well on a complex scenario, but low performance on some datasets suggests that logical metonymy is still a challenging phenomenon for computational modeling. |
Copied to clipboard
| Challenge: | a growing body of theoretical work on narrative has been focused on the field of natural language processing . this position paper aims to provide a unifying framework for the computational study of narrative . |
| Approach: | They propose to introduce dominant theoretical frameworks to the NLP community and situate current research within distinct narratological traditions. |
| Outcome: | The proposed framework would allow for new empirical questions and applications in the field of natural language processing. |
Copied to clipboard
| Challenge: | rap is a popular genre in the u.s. and has been used in countries far beyond the uk . linguistic borrowings are especially intriguing in countries such as the eu and europe . |
| Approach: | They manually annotate a lexicon of over 700 borrowings in the French language . they find that there are increases in the proportion of linguistic borrowings, interjections, and Niger-Congo borrowings . |
| Outcome: | The proposed method analyzes a corpus of over 8000 french rap song lyrics and shows that rap borrowings are increasing in prevalence and interjections are decreasing. |
Copied to clipboard
| Challenge: | Existing methods for assessing argument quality in isolation analyze their quality in the absence of context, which affects their accuracy and generalizability. |
| Approach: | They propose a method for scoring argument quality based on contextualization via relevant knowledge that leverages large language models to provide feedback, infer hidden assumptions, supply a similar-quality argument, or give a counter-argument. |
| Outcome: | The proposed method outperforms existing methods across multiple metrics in both in-domain and zero-shot setups. |
Copied to clipboard
| Challenge: | a stylometric toolkit for analysis of Latin literary texts is available for free at www.qcrit.org/stylometry. |
| Approach: | They propose a stylometric toolkit for analysis of Latin literary texts which generates data for a diverse range of literary features and has an intuitive point-and-click interface. |
| Outcome: | The proposed toolkit generates data for a diverse range of literary features and has an intuitive point-and-click interface. |
Copied to clipboard
| Challenge: | Existing methods for identifying and resolving persona knowledge gaps are underexplored. |
| Approach: | They propose a framework that dynamically detects and resolves persona knowledge gaps using intrinsic uncertainty quantification and feedback-driven refinement. |
| Outcome: | The proposed framework detects and resolves persona knowledge gaps using intrinsic uncertainty quantification and feedback-driven refinement on two real-world datasets: CCPE-M for preferential movie recommendations and ESConv for mental health support. |
Copied to clipboard
| Challenge: | wikiHow guides for specific target groups reflect disparate social norms and subtle stereotypes, a new study shows . wikihow guides are subject to subtle biases, and we aim to raise awareness of these inequalities in future work. |
| Approach: | They investigate the extent to which how-to guides from wikiHow differ in practice depending on intended audience. |
| Outcome: | The findings show that how-to guides from wikiHow differ in practice depending on the intended audience. |
Copied to clipboard
| Challenge: | Detecting online abusive language in social media messages is gaining increasing attention from scholars and stakeholders. |
| Approach: | They propose a hybrid approach with deep learning and a multilingual lexicon to cross-domain and cross-lingual detection of abusive content. |
| Outcome: | The proposed system can detect abusive content across domains and languages using a multilingual lexicon and a domain-independent lexical. |
Copied to clipboard
| Challenge: | Existing dictionaries do not capture the full range of polysemous and homonymous words corresponding to different signs across contexts. |
| Approach: | They analyze 1,404 word use–to–sign ID mappings from German and German Sign Language . they identify three correspondence types: Type 1 (one-to-many), Type 2 (many-to-1), and Type 3 (one to one) |
| Outcome: | The proposed method outperforms existing methods using Exact Match and Semantic Similarity. |
Copied to clipboard
| Challenge: | Existing argumentation datasets have allowed only limited assessment of "user" traits because information on background of users is generally unavailable. |
| Approach: | They present a dataset of 78,376 debates generated over a 10-year period along with surprisingly comprehensive participant profiles. |
| Outcome: | The proposed dataset includes 78,376 debates generated over a 10-year period along with comprehensive participant profiles. |
Copied to clipboard
| Challenge: | Affective responses to music are highly personal, but it's difficult to measure marginal effects of these variables . a study of 403M listener comments on a social music platform in china aims to address this gap . |
| Approach: | They propose to measure affective responses to music from 403M listener comments on a Chinese social music platform. |
| Outcome: | The proposed method identifies musical, lyrical, contextual, demographic, and mental health effects that drive listener affective responses from over 403M listener comments on a Chinese social music platform. |
Copied to clipboard
| Challenge: | Existing methods to protect the identity and privacy of online authorship are lacking supervision data for diverse authorship and domains. |
| Approach: | They propose an unsupervised inference-time approach to authorship obfuscation that uses a user-controlled, inference time algorithm to oblige the authorship. |
| Outcome: | The proposed method outperforms state-of-the-art methods while performing competitively against a propriety model two orders of magnitudes larger. |
Copied to clipboard
| Challenge: | linguistic and social aspects of code-switching are not discussed in the literature in linguistics. |
| Approach: | They propose to examine linguistic and social aspects of code-switching across a wide range of languages in a survey of the literature in linguistics and language technologies. |
| Outcome: | The proposed framework aims to increase the clarity and depth of computational investigations of C-S and bridge the fields so that they might be mutually reinforcing. |
Copied to clipboard
| Challenge: | Citation context analysis (CCA) is an important task in natural language processing that studies how and why scholars discuss each other’s work. |
| Approach: | They propose to use a dataset of 12.6K citation contexts from 1.2K computational linguistics papers to model three important CCA phenomena. |
| Outcome: | The proposed dataset contains 12.6K citation contexts from 1.2K computational linguistics papers and can model these phenomena. |
Copied to clipboard
| Challenge: | In this paper, we analyze memes as a form of language subject to the same kinds of sociolinguistic variation as other modalities, such as written language and speech. |
| Approach: | They propose a computational pipeline to cluster memes into templates and semantic variables, taking advantage of their multimodal structure to learn meme semantics from an unstructured dataset. |
| Outcome: | The proposed method uses 3.8M images from a reddit meme database to analyze linguistic variation in memes. |
Copied to clipboard
| Challenge: | Recent studies show that cognitive manifestations of future dementia may appear as early as 18 years prior to clinical diagnosis . lack of clear diagnosis and prognosis, possibly for an Alzheimer's type, is a major limitation of current methods for identifying dementia-specific cognitive markers. |
| Approach: | They propose to interrogate neural LMs trained on participants with and without dementia by manipulating lexical frequency. |
| Outcome: | The proposed model improves upon the current state-of-the-art for models trained on transcripts of speech produced by healthy participants and those with dementia. |
Copied to clipboard
| Challenge: | Existing methods to detect online abuse focus on the more explicit forms of abuse . existing methods focus on detecting subtler forms of online abuse leaving them unnoticed . |
| Approach: | They propose a task to detect unpalatable questions using reddit data to implement a context-aware dataset and implement 'learning models' they hope future research will address subtle forms of abuse since harm passes unnoticed through existing detection systems. |
| Outcome: | The proposed task is based on a dataset of reddit users and a conversational context. |
Copied to clipboard
| Challenge: | Existing literature on Arabic sentiment analysis is limited, compared to high-resourced languages such as English and French. |
| Approach: | They present a systematic review of existing literature on Arabic sentiment analysis focusing on research utilizing deep learning. |
| Outcome: | The proposed methods highlight gaps in the literature on Arabic sentiment analysis and outline promising directions for future research. |
Copied to clipboard
| Challenge: | a new method to normalize orthographic variations of historical documents is needed for digital humanities and diachronic studies. |
| Approach: | They propose to normalize orthographic wordforms found in Middle French archives . authors say it improves accuracy and accuracy over a strong baseline . |
| Outcome: | The proposed methods normalize orthographic variations of historical documents without modernizing them. |
Copied to clipboard
| Challenge: | a recent study has found that the disclosure of sexual abuse has positive psychological im- pacts. |
| Approach: | They propose to aggregate personal experiences of sexual harassment from Twitter posts to facilitate a better understanding of social media constructs and bring about social change. |
| Outcome: | The proposed model is compared with state-of-the-art models and is based on a three part Twitter-Specific Social Media Language Model. |
Copied to clipboard
| Challenge: | Imageability and concreteness are psycholinguistic properties that link visual and semantic spaces. |
| Approach: | They propose an unsupervised measure that quantifies sharpness of peaks in an image-caption dataset. |
| Outcome: | The proposed method is more robust than existing methods and predicts these properties for classification. |
Copied to clipboard
| Challenge: | Critical toponymy studies the dynamics of power, capital, and resistance through place names and the sites to which they refer. |
| Approach: | They propose a model that measures how cultural and economic capital shape the ways in which people refer to places through an annotated dataset of Airbnb listings in New York City. |
| Outcome: | The proposed model can identify important discourse categories integral to the characterization of place. |
Copied to clipboard
| Challenge: | Suicide is a global problem, with one suicide case for every 100 deaths worldwide . social networking sites are an essential forum for communication and information sharing . |
| Approach: | This paper compares natural language processing to suicidal ideation detection and risk assessment . it urges better intention understanding for reliable suicide risk assessment with computational methods . |
| Outcome: | This paper compares the performance of natural language processing to suicidal ideation detection and risk assessment tasks. |
Copied to clipboard
| Challenge: | a new framework for studying political polarization in social media is needed to understand how group divisions manifest in language. |
| Approach: | They propose to cluster tweet embeddings to uncover four dimensions of political polarization in social media . their results apply existing lexical methods to analyze 4.4M tweets on 21 mass shootings . |
| Outcome: | The proposed framework generates more cohesive topics than traditional models. |
Copied to clipboard
| Challenge: | This paper examines the utility and timeliness of the Hong Kong Protest News Dataset . it sheds light on whether depth and/or manner of reporting changed over time . |
| Approach: | They use the Hong Kong Protest News Dataset to investigate synchronic news characterisations of protests in Hong Kong between 1998 and 2020. |
| Outcome: | The dataset sheds light on whether depth and/or manner of reporting changed over time, and if so, in what ways, or in response to what. |
Copied to clipboard
| Challenge: | Indigenous peoples are increasingly unable to go on without speech and language technologies, says a researcher . a postcolonial approach to computational methods for supporting language vitality is needed, says the researcher - lil'watul Lorna Williams . |
| Approach: | They propose to examine colonising discourses in speech and language technology and propose a postcolonial approach to computational methods for supporting language vitality. |
| Outcome: | The paper reviews colonising discourses in speech and language technology and suggests new ways of working with Indigenous communities. |
Copied to clipboard
| Challenge: | Existing methods to assess lexical complexity are used to evaluate the difficulty of vocabulary for language learners. |
| Approach: | They propose to use pre-trained language models to assess the complexity of a word based on its context. |
| Outcome: | The proposed method outperforms the best systems in SemEval-2021. |
Copied to clipboard
| Challenge: | enhancing the quality of online public discourse requires promoting foundational human virtues, such as “intellectual humility” (IH) . discourse on social media rewards forgetting our virtuous selves, embedding users within echo chambers and causing negative affect towards those who hold different beliefs. |
| Approach: | They propose to use a codebook to measure "intellectual humility" they manually validated the codebook and used it to develop LLM-based models . |
| Outcome: | The proposed model achieves a Macro-F1 score of 0.64 across labels and 0.70 when predicting IH/IA/Neutral at the coarse level. |
Copied to clipboard
| Challenge: | a recent paper argues that current publications foster a gap between adoption and understanding of models . it also makes it easier to meet publication demands with method papers, argues the paper . |
| Approach: | They argue that current NLP publication models foster a gap between adoption and understanding of models . they argue that everlarger models make it harder to explain how our methods work . |
| Outcome: | The authors argue that current publications foster a gap between adoption and understanding of models . they argue that the rise of everlarger models makes it harder to explain how our methods work . |
Copied to clipboard
| Challenge: | Existing methods for protoword reconstruction are limited to a few languages. |
| Approach: | They propose a new database of cognate words and etymons for the five main Romance languages and apply machine learning to it. |
| Outcome: | The proposed model achieves 90% accuracy in predicting protowords for Romance languages, surpassing state-of-the-art models and features. |
Copied to clipboard
| Challenge: | Bantu languages are still computationally under-resourced due to their complex grammatical structure . morphological analyzers, text generation tools and a morphology analyzer are among the tools used . |
| Approach: | They propose a syntactic and semantic method to disambiguate among singular nouns . they use the nearest neighbors of a query word as semantic generalizations based on Runyankore . |
| Outcome: | The proposed method improves accuracy in three Bantu languages compared to using only the syntactic or semantic approach. |
Copied to clipboard
| Challenge: | a computational approach to measure metaphorical language is based on immigration discourse on social media. |
| Approach: | They propose a computational approach that leverages word-level and document-level signals to measure metaphor with respect to immigration discourse on social media. |
| Outcome: | The proposed method measures metaphorical language in immigration discourse on social media. |
Copied to clipboard
| Challenge: | standardized questionnaires are essential tools for mental health screening, but computational approaches bypass these tools in favor of black-box classification. |
| Approach: | They propose a questionnaire-guided screening framework that bridges psychological practice and computational methods through adaptive Retrieval-Augmented Generation. |
| Outcome: | The proposed framework matches or outperforms state-of-the-art performance on Reddit-based benchmarks and extends to self-harm screening. |
Copied to clipboard
| Challenge: | Mental health disorders are a major economic burden for society and are projected to rise to a staggering US $6 trillion by 2030. |
| Approach: | They propose to use memes to identify fine-grained depression symptoms from memes . they benchmark RESTORE on 20 strong monomodal and multimodal methods . |
| Outcome: | The proposed method can predict fine-grained depression symptoms better than existing models that overlook implicit connections between visual and textual elements of a meme. |
Copied to clipboard
| Challenge: | Drug-drug interactions arise when multiple drugs are administered concurrently. |
| Approach: | They propose a pairwise knowledge-augmented generative method for DDIE text generation that integrates biological functions from a knowledge set into a language model. |
| Outcome: | The proposed method outperforms existing methods in DDIE text generation on two professional datasets. |
Copied to clipboard
| Challenge: | Existing methods for measuring Lexical Semantic Change are lacking historical benchmarks. |
| Approach: | They propose a three-stage general-purpose evaluation framework that simulates theory-driven LSC using In-Context Learning and a lexical database. |
| Outcome: | The proposed framework evaluates the sensitivity of computational methods to synthetic change and their suitability for detecting change in specific dimensions and domains. |
Copied to clipboard
| Challenge: | With the introduction of new privacy regulations, disclosures made by the same organization are not always the same in different languages. |
| Approach: | They propose a language annotation scheme to capture nuances of two new privacy regulations, namely the EU’s GDPR and California’s CCPA/CPRA. |
| Outcome: | The proposed method captures the nuances of two new privacy regulations and compares them to a corpus of 64 privacy policies in English and 91 in German with manual annotations for 8K and 19K fine-grained data practices. |
Copied to clipboard
| Challenge: | Existing methods for argument quality assessment do not consider multi-perspective evaluation due to subjective nature of arguments. |
| Approach: | They propose a multi-persona framework for argument quality assessment that simulates diverse evaluator perspectives through large language models. |
| Outcome: | The proposed framework outperforms baselines while providing comprehensive multi-perspective rationales on IBM-Rank-30k and IBM-ArgQ-5.3kArgs datasets. |
Copied to clipboard
| Challenge: | Existing computational methods for framing analysis are limited . a lack of a comprehensive understanding of framability is limiting the research . |
| Approach: | They propose to combine existing approaches to analyze large-scale datasets using computational methods. |
| Outcome: | The proposed methods will help scholars better understand how frames are being explored computationally, the authors argue . |
Copied to clipboard
| Challenge: | Annotated corpus of English cooking recipe procedures with domain-specific linguistic and semantic structure. |
| Approach: | They annotate a corpus of English cooking recipe procedures with domain-specific linguistic and semantic structure and then use a flow graph to represent the sequence of steps. |
| Outcome: | The proposed methods achieve 71.1 to 87.5 F1 in the cooking domain and a flow graph achieves similarity to those used in Japanese recipes. |
Copied to clipboard
| Challenge: | Existing studies show that word choice is driven by demographics within the United States. |
| Approach: | They develop computational methods to study word choice within a sociolinguistic lexical variable . they use two variables to test for attitudes towards sexuality and gender in the u.s. |
| Outcome: | The proposed methods allow us to examine attitudes towards sexuality and gender in the United States through two lexical variables. |
Copied to clipboard
| Challenge: | Language model-based agents can be used to conduct and support data-driven science, but evaluating them on open-ended tasks is challenging due to multiple valid approaches, partially correct steps, and different ways to express the same decisions. |
| Approach: | They propose a benchmark to automatically evaluate agents’ multifaceted approaches to open-ended research questions. |
| Outcome: | BLADE evaluates agents’ multifaceted approaches to open-ended research questions using data from 12 datasets and research questions drawn from existing scientific literature. |
Copied to clipboard
| Challenge: | Existing computational methods for DDI prediction fail to capture interactions for new drugs due to the lack of knowledge. |
| Approach: | They propose a problem setup as zero-shot DDI prediction that deals with the case of new drugs by using textual information from online databases. |
| Outcome: | The proposed method improves on several settings including zero-shot and few-shot DDI prediction and the selected texts are semantically relevant. |
Copied to clipboard
| Challenge: | a recent study identifies unreliable narrators, i.e. those who unintentionally misrepresent information . authors propose using computational methods to identify unredependable narrators . adbrei: readers implicitly question the reliability of the nrator . |
| Approach: | They propose using computational methods to identify unreliable narrators . they use literary theory to define different types of unredependable narrators . |
| Outcome: | The proposed method can identify unreliable narrators on real-world text data. |
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have remarkable out-of-domain generalizability to novel optimization tasks. |
| Approach: | They propose a series of instruction-tuned LLMs for molecule optimization that outperform state-of-the-art instruction-based LLM models. |
| Outcome: | mathttMuMOInstruct outperforms state-of-the-art LLMs on 5 in-domain and 5 out-of domain tasks. |
Copied to clipboard
| Challenge: | Existing methods to determine the semantic relatedness between compounds and constituents have applied a synchronic perspective, but this study examines what diachronic changes in contexts and semantic topics reveal about the compounds’ present-day compositionality. |
| Approach: | They propose to use two diachronic vector spaces to model compositional patterns between compounds with low and high present-day compositionality. |
| Outcome: | The proposed model performs on par with co-occurrence space and captures similar information. |
Copied to clipboard
| Challenge: | Personal name compounds (PNCs) are compositions that refer to a person, such as Willkommens-Merkel ('Welcome-Meerkel') and a personal name such as Merkel. |
| Approach: | They propose to model 321 personal name compounds and their corresponding full names at discourse level and compare two approaches to assess whether a PNC is more positively or negatively evaluative . they further enrich data with personal, domain-specific, and extra-linguistic information and perform regression analyses revealing that factors including compound and modifier valence, domain, and political party membership influence how a pnc is evaluated. |
| Outcome: | The proposed model shows that the PNCs are perceived as more positively or negatively than their full name and that they are perceived to be more positive or negative. |
Copied to clipboard
| Challenge: | Comics offer compelling yet under-explored domain for computational narrative analysis . authors highlight potential of comics for narrative-driven, multimodal data analysis based on novel comics . |
| Approach: | They propose a dataset of scene-level narrative arcs derived from comic books . they highlight their potential to inform broader research on multimodal storytelling . |
| Outcome: | The dataset provides an initial benchmark that future studies can build upon. |